Repository navigation
fix(agent-core-v2): fold UserPromptSubmit hook output into the prompt message - #4081
Conversation
…tent parts
UserPromptSubmit hook messages now merge into the user prompt message as
meta-marked text content parts (contentType text/xml, source 'user prompt
submit hook') at content position 0 instead of a standalone hook_result
user message, so one prompt always yields exactly one user message and
turn pairing can no longer shift. Structured hook JSON without a message
field (e.g. {"continue":true}) no longer injects anything; plain-text
stdout still falls through. turn.prompt records, TurnStarted prompts,
fold turn openings, and steer content keys all strip hook parts; existing
wires with standalone hook_result messages keep their current behavior.
🦋 Changeset detectedLatest commit: 1941130 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
commit: |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ec03c39a97
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| ...ctx.hookParts, | ||
| ...gateImageFormatParts(promptMessage.content, this.profile.getModelProviderType()), |
There was a problem hiding this comment.
Exclude injected parts from active prompt projections
When a UserPromptSubmit hook returns text, this prepends a meta-marked part to the active prompt. The /sessions/{session_id}/prompts projection still forwards that active message verbatim (and, for promptWithSkills, uses slice(skillActivations.length)), while projectPromptContentParts drops the meta. Consequently REST/transcript prompt consumers display the internal <hook_result> XML—and bundled prompts also expose a rendered skill block—instead of the caller’s submitted content; they have no metadata left to filter it. Strip hook and skill parts before those projections.
AGENTS.md reference: AGENTS.md:L100-L100
Useful? React with 👍 / 👎.
| const text = annotateBundledSkillParts(input, bundledActivations) | ||
| .filter((part): part is TextPart => part.type === 'text' && !isSkillActivationPart(part)) | ||
| .map((part) => part.text) |
There was a problem hiding this comment.
Keep hook output out of TurnStarted prompts
When a UserPromptSubmit hook returns output, runPromptGate prepends it to the prompt and this helper removes only skill parts. TurnStarted.prompt therefore contains the hook's <hook_result> alongside the caller input; AgentExternalHooksService.notifyTurnStarted forwards that value as the next TurnStarted hook's inputData.prompt. Hook scripts that inspect, log, or match submitted prompts now receive their own prior hook output rather than the caller prompt. Filter isUserPromptSubmitHookPart here as well.
AGENTS.md reference: AGENTS.md:L100-L100
Useful? React with 👍 / 👎.
| prompt: isDisplayablePromptOrigin(prompt.origin) | ||
| ? turnPromptText(prompt.message.content, prompt.origin) | ||
| : undefined, | ||
| promptContent: isDisplayablePromptOrigin(prompt.origin) ? prompt.message.content : undefined, |
There was a problem hiding this comment.
Strip promptContent from public turn events
For every displayable prompt, this adds the full internal ContentPart[] to turn.started. Both public bridges strip only promptAttachments (packages/node-sdk/src/v2/event-mapper.ts and packages/kap-server/src/transport/ws/v1/sessionEventBroadcaster.ts), so promptContent is serialized to SDK and WebSocket clients even though it is absent from turnStartedEventSchema and the established wire field set. This duplicates potentially large data-URL media in durable event traffic and exposes internal hook markup and metadata; remove promptContent alongside promptAttachments in those bridges.
AGENTS.md reference: AGENTS.md:L112-L112
Useful? React with 👍 / 👎.
| ...ctx.hookParts, | ||
| ...gateImageFormatParts(promptMessage.content, this.profile.getModelProviderType()), |
There was a problem hiding this comment.
Hide hook parts from message-history responses
When a UserPromptSubmit hook emits text, this makes its <hook_result> a part of the durable user message. buildProtocolContent maps every user part to public MessageContent while dropping meta, and both GET /sessions/{session_id}/messages and the session message pagination call toProtocolMessage. Those responses therefore expose raw hook protocol XML inside the user's message with no marker left for clients to distinguish it, whereas the old standalone message carried origin: hook_result. Filter or separately project these parts before converting message history to the REST contract.
AGENTS.md reference: AGENTS.md:L112-L112
Useful? React with 👍 / 👎.
57d0eed to
3f77226
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3f7722633c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| prompt: isDisplayablePromptOrigin(prompt.origin) | ||
| ? turnPromptText(prompt.message.content, prompt.origin) | ||
| : undefined, | ||
| promptContent: isDisplayablePromptOrigin(prompt.origin) ? prompt.message.content : undefined, |
There was a problem hiding this comment.
Keep hook text out of public turn prompts
When a UserPromptSubmit hook returns text, this computes TurnStarted.prompt from the injected hook part as well as the caller input. Both public bridges then remove promptContent (packages/node-sdk/src/v2/event-mapper.ts:69-75 and packages/kap-server/src/transport/ws/v1/sessionEventBroadcaster.ts:966-972), leaving SDK and WebSocket consumers with no metadata to distinguish the internal <hook_result> XML from submitted text. Clients that display or log turn.started.prompt as the user's input will now expose hook output; filter hook parts before assigning the public prompt.
AGENTS.md reference: AGENTS.md:L110-L112
Useful? React with 👍 / 👎.
63a91ea to
bf5f5fe
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bf5f5fe912
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| content: [ | ||
| ...message.content.filter( | ||
| (part) => isUserPromptSubmitHookPart(part) || isSkillActivationPart(part), | ||
| ), | ||
| { type: 'text', text }, |
There was a problem hiding this comment.
Avoid duplicating preserved parts during compaction
When full compaction truncates or splits a long prompt containing a UserPromptSubmit part (or a bundled skill block), extractText(message.content) has already included that text in the truncated payload. This new filter then retains the same marked part before appending the payload, so the compacted head/boundary message sends the hook or skill content to the model twice and can exceed the intended token budget; exclude retained parts from the replacement text or avoid preserving them in the prefix.
Useful? React with 👍 / 👎.
bf5f5fe to
63bdfcb
Compare
… and handle them across client surfaces
…onal slicing
Bundled skill blocks are now constructed with meta { source: 'skill
activation' }, matching the hook-part convention. Every former
content.slice(skillActivations.length) site — turn prompt text, steer
merge/echo, session and fork titles, transcript fold, TUI replay — now
annotates then filters by meta, so prompt content no longer depends on
block positions and the loop needs no hook-part knowledge. A shared
annotate step keeps legacy wires working: leading unmarked blocks (one
per bundled activation) are marked on read before filtering.
…mpt at construction
63bdfcb to
1941130
Compare
Step 11, the tail of the sync: 13 upstream commits over 177 files. The step lands the WaitFor steering rework, the workspace-trust summary, the sticky user messages, and the completion-budget cleanup. Upstream's changes in this step, all absorbed: - interrupt tower wake turns on user input and harden tower operations (MoonshotAI#3998) - cap WaitFor at 90s and let new input end the wait (MoonshotAI#4061) - summarize workspace trust with configuration sources (MoonshotAI#4056) - trust the workspace via KIMI_CODE_TRUST_WORKSPACE (MoonshotAI#4059) - add sticky user messages in fullscreen mode (MoonshotAI#4031) - omit the completion token cap unless explicitly configured (MoonshotAI#4091) - clear inherited cron tasks at the fork record (MoonshotAI#4083) - fold UserPromptSubmit hook output into the prompt message (MoonshotAI#4081) - remove the interruption reminder on undo of its turn (MoonshotAI#4076) - publish permission mode changes on agent.status.updated (MoonshotAI#4057) - stop NotifyUser nudges in hosts without the update panel (MoonshotAI#4054) - add regression and user-impact review to the PR workflow (MoonshotAI#4023) - sync the 2.1.1 changelog into the release notes (MoonshotAI#4018) Fork behaviour that had to survive: - the message-dispatch controller split: upstream's new steering path (the steering set, the queue-wide steer, the settle callback, the transcript rollback on a failed steer) was ported into MessageDispatchController, and KimiTUI keeps delegating to it - the i18n layer: the rewritten trust prompt goes through `t()` with new en/zh entries, and the Code Review Rules section lands in DEVELOP.md (the fork's renamed AGENTS.md) - the fork's node-local → host rename, its sixel graphics work in pi-tui, and its Windows-safe test cleanup stay in place - the fork's tool list: the recorded tools-snapshot hashes keep the fork's own tool set - the fork's path-image ingestion no longer defers the media-free submit path: extractMediaAttachments returns synchronously when the text carries no image file path, so a plain prompt reaches dispatch in the same tick again (upstream's steering tests assert exactly that) Verified: typecheck clean across every package and app; full suite 17879 passed / 0 failed / 81 skipped / 1 todo, plus kimi-inspect's own 115 tests; lint 0 errors.
Requirement or Bug
Resolve #4032
Supersedes #4078(评审历史见该 PR)。
Bug Reproduction Steps
See linked issue(2.1.0/2.1.1 可复现:配置
UserPromptSubmit钩子的会话里,轮次整体错位一位,最后一条消息看起来永远没有回复,刷新后恢复)。Root Cause
引擎在
UserPromptSubmit钩子返回文本后,把它作为一条独立的 user 角色消息(origin: hook_result)追加到真实用户消息之前,上下文序列变成[hook 消息][用户消息][回复],user 单元翻倍。渲染层的轮次配对依赖「一个 user 单元配一段 assistant 输出」的不变量,任何不消费origin的配对路径都会因多出的 user 单元而整体错位一位。排查确认数据层(wire、fold、live store、投影)完整且对齐,错位只来自不消费 hook origin 的渲染路径。这是根本修复而非绕过:让「一条 prompt 一条 user 消息」在结构上成立——钩子注入文本以带 meta 标记的 content part 并入 user prompt 消息,不再产生第二条 user 消息,轮次配对在结构上不存在错位可能,不再依赖每条渲染路径自觉遵守 origin 约定。
Code Changes
三个提交,按主题分六块:
1. content part 增加 meta 能力(
agent-core-v2+transcript)wire schema(
historySchema.ts)以z.record形态接受meta(通用字典,前向兼容);packages/transcript的HistoryContentParttext 变体同步(browser-safe 纯类型)。meta 不下发 LLM provider(adapter 只读type/text)。2. 钩子注入合并进 prompt 消息(gate 单次落盘)
PromptSubmitContext增加hookParts可变插槽(钩子的唯一输出通道,与既有block字段同模式);hook part 固定形状{ type:'text', text:'<hook_result …>…</hook_result>', meta:{ contentType:'text/xml', source:'user prompt submit hook' } },多钩子按返回顺序各成一个 part。message字段(如{"continue":true})不再注入任何内容(此前原文注入,每轮都有噪声);纯文本 stdout 仍兜底注入。3. Turn 层记录保留完整数据(vis/replay 可见)
turn.promptwire 记录含完整 hook 内容(持久化、进上下文);TurnStarted事件不落 wire,其prompt在构造点按 meta 剥离 hook part——公开客户端拿到干净文本,持久化与上下文语义不受影响。4. 展示层可见渲染 / 干净摘要
hookmarker(载荷与 livehook.resultmarker 同形),轮 prompt 提取/steer 匹配仍按剥离后内容加工,不产生 0 步轮组;turn.started.prompt(构造点已按 meta 剥离,投影不做任何字符串加工,用户手敲相同 wrapper 开头原样保留);TurnBegin.user_input剥 hook part,hook 正文以 StepBegin 后的 ContentPart 可见渲染(webview reducer 要求 step 存在,否则丢弃);/export-md与 VS Code/export的 markdown、VS Code/import <session>的会话文本:均剥 hook part,内部协议 XML 不再进入导出文件或被再注入其他会话;lastPrompt重算(session 列表摘要)剥 hook part,与 title/fork 路径对齐;5. skill bundled 块统一 meta 化,消除位置切片
skill 块与 hook part 一样在构造时携带
meta: { source: 'skill activation', activationId }(fallback 补标时按 origin 顺序对应一并回填 activationId,新老 wire 的 skill part 均自描述)。原先全部content.slice(skillActivations.length)位置切片点——轮 prompt 文本、steer 合并与 echo、会话标题与 fork 标题、transcript fold、TUI replay——统一改为「先 annotate 再按 meta 过滤」:annotate 是集中的向后兼容层,老 wire 里无标记的前 N 个块(每个 bundled activation 一个)在读取时补标,之后与有标记的新 wire 走同一条过滤路径。prompt 内容不再依赖块位置,loop 也无需识别 hook part。Behavior Changes and Affected Users
UserPromptSubmit注入文本的上下文形态origin: hook_result),位于真实 prompt 之前message的 JSON(如{"continue":true}){"message": "..."}turn.prompt记录 /turn.started事件的 promptturn.prompt记录含完整 hook 内容(持久化、进上下文);turn.started事件的 prompt 在构造点按 meta 剥离 hook part(事件不落 wire,持久化不受影响)turn.started.prompt的客户端——ACP、desktop、SDK——拿到的是干净文本)origin.skillActivations.length位置识别meta: { source: 'skill activation', activationId }(含 steer 合并消息;老 wire 读取时由 fallback 回填)hook.result事件(一张卡)hook.result事件(各成一张卡),与重开后 transcript 的逐 part marker 一致受影响模块与测试覆盖(测试数净增为 0,均为扩展既有用例):
agent-core-v2gate / 钩子服务):promptService.test.ts的 gate 用例钉住单条 user 消息、content[0] meta 形状、turn.prompt记录含 hook 内容、TurnStarted.prompt构造点剥离为干净文本、skill 块不进 prompt 文本;runHook/userPrompt):runner.test.ts钉住无messageJSON(含{"continue":true})不注入、纯文本 stdout 产出带约定 meta 的 part;origin.ts/ steer 合并):promptService.test.ts的 steer/queue 用例与machine.test.ts钉住标记与未标记块两种形态的合并/echo 结果;transcript、kap-server):layers.test.ts钉住 hook marker 可见、bundled 两形态一致、steer content-key 匹配;transcript.test.ts钉住轮标题 meta 级剥离;apps/kimi-code):message-replay.test.ts、kimi-tui-message-flow.test.ts、export-markdown.test.ts钉住 hook 条目可见、prompt/导出文本干净;apps/vscode):replay-adapter.test.ts钉住 TurnBegin 干净、hook 经 StepBegin 后 ContentPart 可见;agent-core-v2):undo.test.ts钉住锚点与 pending 两分支剥 hook part;Checklist
/approve).gen-changesetsskill, or this PR needs no changeset.gen-docsskill, or this PR needs no doc update.(state-manifest.d.ts已随类型同步重新生成)